Papers with preprocessing step

19 papers
Neural Mention Detection (2020.lrec-1)

Copied to clipboard

Challenge: Mention detection is an important preprocessing step for downstream applications such as NER and coreference resolution.
Approach: They propose and compare three approaches to mention detection using ELMO embeddings and a biaffine classifier.
Outcome: The proposed model outperforms state-of-the-art models on the GENIA corpora and improves on mention recall.
CoQAR: Question Rewriting on CoQA (2022.lrec-1)

Copied to clipboard

Challenge: Existing systems that ask questions in a conversational context may have contextual dependencies that make the understanding difficult.
Approach: They propose to rewrite questions into an out-of-context form to facilitate understanding . they propose to use this form to train and evaluate conversational question answering models .
Outcome: The proposed model can be used in the supervised learning of three tasks: question paraphrasing, question rewriting and conversational question answering.
Multilingual Email Zoning (2021.eacl-srw)

Copied to clipboard

Challenge: Existing literature on email zoning is mainly limited to English . however, it is possible to discern a level of formal organization in the way most emails are formed.
Approach: They propose a multilingual email zoning benchmark based on a language agnostic sentence encoder and a new model that uses a biLSTM with a CRF to classify each sentence into an email zone.
Outcome: The proposed model is competitive with current English benchmarks and reached state-of-the-art performance in English.
Discourse-Based Sentence Splitting (2021.findings-emnlp)

Copied to clipboard

Challenge: Sentence splitting is a key component of sentence simplification and has been shown to help human comprehension.
Approach: They propose to use a discourse connective to generate a sentence that is shorter than the input text.
Outcome: The proposed models outperform end-to-end models in learning the various ways of expressing a discourse relation but generate text that is less grammatical.
Goodwill Hunting: Analyzing and Repurposing Off-the-Shelf Named Entity Linking Systems (2021.naacl-industry)

Copied to clipboard

Challenge: Named entity linking (NEL) is a preprocessing step in commercial systems . a small organization or individual could use an off-the-shelf system to accomplish the same objectives .
Approach: They examine how to repurpose off-the-shelf NEL systems to correct sport-related errors.
Outcome: The proposed model can improve sports question-answering accuracy by 25% . the proposed model is based on the best available model .
A Simple Approach for Handling Out-of-Vocabulary Identifiers in Deep Learning for Source Code (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods to handle out-of-vocabulary identifiers are not suitable for source code processing.
Approach: They propose a method to handle out-of-vocabulary identifiers by identifies anonymization . they show that the method significantly improves the performance of the Transformer .
Outcome: The proposed method significantly improves the performance of the Transformer in two code processing tasks.
Increasing In-Class Similarity by Retrofitting Embeddings with Demographic Information (D18-1)

Copied to clipboard

Challenge: a new method for text classification ignores strong non-linguistic similarities like homophily . authors are typically represented via their linguistic profiles, i.e. information avail-able in the text .
Approach: They use homophily cues to retrofit text-based author representations with non-linguistic information and introduce a trade-off parameter.
Outcome: The proposed method improves on two author-attribute prediction tasks with large labels.
Multilingual Normalization of Temporal Expressions with Masked Language Models (2023.eacl-main)

Copied to clipboard

Challenge: Existing methods for normalizing temporal expressions are rule-based, which severely limits the applicability in multilingual settings.
Approach: They propose a neural method for normalizing temporal expressions based on masked language modeling and a slot-based prediction scheme for context-independent representations.
Outcome: The proposed method outperforms existing rule-based methods in many languages and in particular, for low-resource languages with performance improvements of up to 33 F1 on average compared to the state of the art.
Evaluating Historical Text Normalization Systems: How Well Do They Generalize? (N18-2)

Copied to clipboard

Challenge: Historical text normalization systems aim to convert historical wordforms to their modern equivalents . many of these systems have been developed and tested on a single language .
Approach: They propose to use a nave baseline system to evaluate historical text normalization systems . they show that the models generalize well to unseen words in tests on five languages .
Outcome: The proposed models generalize well to unseen words on five languages, but provide no clear benefit over the nave baseline.
DR-BiLSTM: Dependent Reading Bidirectional LSTM for Natural Language Inference (N18-1)

Copied to clipboard

Challenge: Existing approaches to natural language inference rely on simple reading mechanisms for independent encoding of the premise and hypothesis.
Approach: They propose a novel bidirectional dependent reading network to efficiently model the relationship between a premise and a hypothesis during encoding and inference.
Outcome: The proposed model outperforms existing methods by a considerable margin on the Stanford Natural Language Inference (SNLI) dataset.
Fast WordPiece Tokenization (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for tokenization of text are not efficient, but they are based on Aho-Corasick's algorithm.
Approach: They propose an efficient algorithm for WordPiece tokenization using a longest-match-first strategy . they propose an algorithm whose tokenization complexity is strictly O(n)
Outcome: The proposed method is 8.2x faster than HuggingFace Tokenizers and 5.1x faster on average for general text tokenization.
Subword Segmental Machine Translation: Unifying Segmentation and Target Sentence Generation (2023.findings-acl)

Copied to clipboard

Challenge: Subword segmenters are used in neural machine translation, but are not used in high-resource settings.
Approach: They propose a subword segmental machine translation (SSMT) that unifies subword and MT in a single trainable model.
Outcome: The proposed model improves chrF scores for morphologically rich agglutinative languages and is more robust on a test set constructed for evaluating morphology generalisations.
Parsivar: A Language Processing Toolkit for Persian (L18-1)

Copied to clipboard

Challenge: a preprocessing step is required to convert text into a standard format for NLP tasks.
Approach: They propose a Persian preprocessing toolkit that performs various kinds of activities . they use a plagiarism detection system to exploit the proposed toolkit .
Outcome: The proposed tool outperforms available Persian preprocessing tools by about 8 percent in terms of F1 . the proposed toolkit performs normalization, space correction, tokenization, stemming, parts of speech tagging and shallow parsing tasks.
From Characters to Words: Hierarchical Pre-trained Language Model for Open-vocabulary Language Understanding (2023.acl-long)

Copied to clipboard

Challenge: Current models for natural language understanding require a preprocessing step to convert raw text into discrete tokens.
Approach: They propose a hierarchical open-vocabulary language model that adopts a shallow Transformer architecture to learn word representations from their characters and a deep inter-word Transformer module that contextualizes each word representation by attending to the entire word sequence.
Outcome: The proposed model outperforms baselines on various downstream tasks and is robust to textual corruption and domain shift.
Harnessing Pre-Trained Neural Networks with Rules for Formality Style Transfer (D19-1)

Copied to clipboard

Challenge: Existing studies normalize informal sentences with rules, but they introduce noise if we use them in a naive way.
Approach: They propose to harness rules into a state-of-the-art neural network that is typically pretrained on massive corpora.
Outcome: The proposed method can be used to generate a state-of-the-art on a small dataset.
Stop Taking Tokenizers for Granted: They Are Core Design Decisions in Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Subword tokenization approaches misalign with linguistic structure and waste capacity across languages and domains.
Approach: They argue for a context-aware framework that integrates tokenizer and model co-design . they argue that tokenization should be treated as a core design problem, not an afterthought .
Outcome: The proposed framework integrates tokenizer and model co-design, guided by linguistic, domain, and deployment considerations.
Subword Segmental Language Modelling for Nguni Languages (2022.findings-emnlp)

Copied to clipboard

Challenge: Subword segmentation is a standard practice in NLP, but is viewed as a preprocessing step for low-resource languages with complex morphologies.
Approach: They propose a subword segmental language model that learns how to segment words while being trained for autoregressive language modelling.
Outcome: The proposed model outperforms existing models on unsupervised morphological segmentation and outperfies standard subword segmenters on all 4 languages.
Curating Datasets for Better Performance with Example Training Dynamics (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to improve data quality but rely on data quantity to improve performance are not effective.
Approach: They propose a method for weighing the relative importance of examples in a dataset based on their Example Training dynamics (ETD) they propose an active learning approach for computing ETD during training rather than as a preprocessing step.
Outcome: The proposed method can be used to improve performance in in-distribution and out-of-distortion testing.
One Model is All You Need: ByT5-Sanskrit, a Unified Model for Sanskrit NLP Tasks (2024.findings-emnlp)

Copied to clipboard

Challenge: Morphologically rich languages are notoriously challenging to process for downstream NLP applications.
Approach: They propose a pretrained model for NLP applications involving the morphologically rich language Sanskrit that outperforms previous models by a considerable margin.
Outcome: The proposed model outperforms tokenized models on established Sanskrit word segmentation tasks and matches the current best lexicon-based model.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations